The LLM compute market — training, fine-tuning, and inference
How is LLM compute demand, revenue, and profit split between training, fine-tuning, and inference, and what does the shift toward inference mean for AI factory builders?
Companies spent $37B on enterprise generative AI in 2025 and Gartner counts $2.59T of total AI spending for 2026, and the mix is shifting decisively from training to inference — inference overtakes training in AI-optimized cloud spend by 2026 and across the wider market by 2029. Fine-tuning is a real but small slice of that spend, agentic workloads are the fastest-growing inference category, and frontier training is multimodal-first even though public accounting of compute-by-modality barely exists. The evidence says an AI factory builder should not sell inference tokens against its own customers, but should become inference-workload-aware as infrastructure — disaggregated prefill and decode, SLO-aware scheduling, and tokens-per-watt are all sellable attributes.
Updated 19 Aug 202628 sources2019–2026Standard19 min read
LLM inference · training economics · fine-tuning · agentic AI · AI factories · GPU cloud